Model card

A multimodal emotion model, measured in the open

Every figure on this page is transcribed from the training notebook shipped with the project — the architecture, dataset counts and the exact scoring rules. Nothing is estimated.

Audio

7.5M

CNN + BiLSTM + RNN over log-Mel spectrograms

Video

33.3M

ResNet3D r3d_18 over 16-frame clips

Fusion

0.2M

MLP over the 568-dim fused vector

From recording to indicator

What a submitted video walks through before a result appears.

  1. 1

    Extract audio

    ffmpeg → 16 kHz mono WAV

  2. 2

    Score speech

    128-band log-Mel → CNN + BiLSTM + RNN

  3. 3

    Score facial motion

    16 frames @ 112 px → ResNet3D r3d_18

  4. 4

    Fuse both signals

    568-dim vector → Fusion MLP

  5. 5

    Map to indicators

    Softmax → rule-based indicator scores

Architecture

What each stage is built from and the signal it consumes.

Audio

CNN + BiLSTM + RNN over log-Mel spectrograms

Layers

  • CNN 64 → 128 → 256 → 512
  • BiLSTM 256 → 512
  • SimpleRNN 512
  • Dense 128 → 64
  • Softmax 8

Specifications

Input
128-band log-Mel, 188 × 128
Sample rate
16 kHz
FFT window
1024
Hop length
256
Pre-emphasis
0.97
Class balancing
SMOTE

Video

ResNet3D r3d_18 over 16-frame clips

Layers

  • ResNet3D r3d_18 backbone (pretrained)
  • Frozen stem + layer1
  • 3D CNN head 256 → 128 → 64
  • Dropout 0.3
  • Softmax 8

Specifications

Input
16 × 112 × 112 RGB clip
Sequence length
16 frames
Frame size
112 px
Parameters
33.3 M
Transfer learning
Kinetics-400 weights

Fusion

MLP over the 568-dim fused vector

Layers

  • Concatenate 568-dim vector
  • Dense 256 → 128 → 64 → 32
  • Dropout 0.3
  • Softmax 8
  • Rule-based indicator layer

Specifications

Input
568-dim fused vector
Composition
MFCC 40 + audio softmax 8 + R3D 512 + video softmax 8
Weight decay
1e-4
LR scheduler
StepLR ×0.5 / 30 epochs

Dataset

Trained and evaluated on paired audio/video recordings of acted emotional speech and song.

RAVDESS

Speech + Song, paired audio/video

Paired samples
2,452
Classes
8
Test split
20%
Validation split
20%

Class distribution

Calm 376
Happy 376
Sad 376
Angry 376
Fearful 376
Disgust 192
Surprised 192
Neutral 188

Indicator scoring

The fusion softmax is turned into four indicators by a transparent rule layer — no second model, no learned weights, so every score can be traced back to the emotion probabilities above.

Possible depression

High sadness together with low happiness. The multiplication means both conditions must hold before the score rises.

score = sad × (1 − happy)

Possible anxiety

Driven by fear, with a small reinforcement from surprise.

score = 0.85 × fearful + 0.15 × surprised

Possible stress

Driven by anger, reinforced by fear.

score = 0.6 × angry + 0.4 × fearful

Possible emotional blunting

A dominant, flat neutral affect across the recording.

score = neutral

No indicator is named unless it scores at least 15%. Below that the emotional pattern is reported as stable.

This result is not a medical diagnosis. Please consult a professional.

Runtime & training

The environment the published metrics were produced on, and the settings each stage was trained with.

Accelerator
NVIDIA Tesla T4 (14.6 GB)
PyTorch
2.10.0 + cu128
TensorFlow
2.20.0
Random seed
42

Hyper-parameters

Audio

Learning rate
3e-4
Batch size
16
Epochs
100 (early stopped)
Dropout
0.2
Weight decay
1e-4
Early stopping
20 epochs

Video

Learning rate
1e-4
Batch size
8
Epochs
25
Dropout
0.3
LR schedule
×0.5 after 3 stale epochs
Backbone
stem + layer1 frozen

Fusion

Learning rate
1e-3
Batch size
32
Epochs
150 (early stopped)
Dropout
0.3
Weight decay
1e-4
Early stopping
25 epochs

See it run on your own recording

The same three networks, scored live on a two-minute video.

Get Started